Chapter 3.1 Normal Positional Embeddings
Normal Positional Embeddings are Simple. Take token Embeddings, add a Learned Matrix which gets better by training called Positional Matrix, add each element by each element and you are done.
Token Embeddings
[ 0.43 0.15 0.89 0.55 0.87 0.66 0.57 0.85 0.64 0.22 0.58 0.33 0.77 0.25 0.10 ] \begin{bmatrix}
0.43 & 0.15 & 0.89 \\
0.55 & 0.87 & 0.66 \\
0.57 & 0.85 & 0.64 \\
0.22 & 0.58 & 0.33 \\
0.77 & 0.25 & 0.10
\end{bmatrix} 0.43 0.55 0.57 0.22 0.77 0.15 0.87 0.85 0.58 0.25 0.89 0.66 0.64 0.33 0.10
Positional Embeddings
[ 0.10 0.20 0.30 0.20 0.30 0.40 0.30 0.40 0.50 0.40 0.50 0.60 0.50 0.60 0.70 ] \begin{bmatrix}
0.10 & 0.20 & 0.30 \\
0.20 & 0.30 & 0.40 \\
0.30 & 0.40 & 0.50 \\
0.40 & 0.50 & 0.60 \\
0.50 & 0.60 & 0.70
\end{bmatrix} 0.10 0.20 0.30 0.40 0.50 0.20 0.30 0.40 0.50 0.60 0.30 0.40 0.50 0.60 0.70
Token + Positional Embeddings
[ 0.43 0.15 0.89 0.55 0.87 0.66 0.57 0.85 0.64 0.22 0.58 0.33 0.77 0.25 0.10 ] + [ 0.10 0.20 0.30 0.20 0.30 0.40 0.30 0.40 0.50 0.40 0.50 0.60 0.50 0.60 0.70 ] = [ 0.53 0.35 1.19 0.75 1.17 1.06 0.87 1.25 1.14 0.62 1.08 0.93 1.27 0.85 0.80 ] \begin{bmatrix}
0.43 & 0.15 & 0.89 \\
0.55 & 0.87 & 0.66 \\
0.57 & 0.85 & 0.64 \\
0.22 & 0.58 & 0.33 \\
0.77 & 0.25 & 0.10
\end{bmatrix}
+
\begin{bmatrix}
0.10 & 0.20 & 0.30 \\
0.20 & 0.30 & 0.40 \\
0.30 & 0.40 & 0.50 \\
0.40 & 0.50 & 0.60 \\
0.50 & 0.60 & 0.70
\end{bmatrix}
=
\begin{bmatrix}
0.53 & 0.35 & 1.19 \\
0.75 & 1.17 & 1.06 \\
0.87 & 1.25 & 1.14 \\
0.62 & 1.08 & 0.93 \\
1.27 & 0.85 & 0.80
\end{bmatrix} 0.43 0.55 0.57 0.22 0.77 0.15 0.87 0.85 0.58 0.25 0.89 0.66 0.64 0.33 0.10 + 0.10 0.20 0.30 0.40 0.50 0.20 0.30 0.40 0.50 0.60 0.30 0.40 0.50 0.60 0.70 = 0.53 0.75 0.87 0.62 1.27 0.35 1.17 1.25 1.08 0.85 1.19 1.06 1.14 0.93 0.80
simple.